In hands-free communication systems, acoustic echo considerably reduces speech quality and intelligibility, especially in reverberant environments. A hybrid acoustic echo cancellation system that combines modulation-domain speech augmentation, Normalized Least Mean Square (NLMS) adaptive filtering, and Room Impulse Response (RIR) simulation is presented in this research. The Image Source Method is used to create realistic acoustic environments and produce RIRs for three different reverberation situations. While STFT-based Wiener filtering minimizes residual echo and background noise while maintaining the intended speech signal, the NLMS adaptive filter suppresses the dominating acoustic echo. Using ERLE, SNRout, SI-SDR, and STOI, the suggested framework is assessed under six typical background noise circumstances with input SNR levels ranging from ?10 dB to 20 dB. Echo suppression, voice quality, and intelligibility all consistently increase in various acoustic situations, according to experimental data.
Introduction
This paper presents a two-stage acoustic echo cancellation (AEC) framework that combines Normalized Least Mean Square (NLMS) adaptive filtering with modulation-domain speech enhancement to improve speech quality in hands-free communication systems. Acoustic echo is a common problem in applications such as voice assistants, smart speakers, and video conferencing, where loudspeaker output is reflected by surrounding surfaces and captured by the microphone. Although the NLMS algorithm effectively removes the primary acoustic echo, residual echo and background noise often remain, degrading speech intelligibility and overall communication quality.
To address this limitation, the proposed system integrates Room Impulse Response (RIR) simulation, NLMS adaptive filtering, and modulation-domain speech enhancement into a unified framework. Realistic acoustic environments are generated using the Image Source Method, which simulates room reverberation. After the dominant echo is removed by the NLMS adaptive filter, the remaining residual echo and background noise are further suppressed using modulation-domain processing based on the Short-Time Fourier Transform (STFT), Noise Power Spectral Density (PSD) estimation, Wiener filtering, and Inverse STFT (ISTFT) reconstruction. This approach preserves the desired speech signal while improving speech quality and intelligibility.
The literature review highlights that adaptive filtering techniques such as LMS, NLMS, and Recursive Least Squares (RLS) are widely used for acoustic echo cancellation, with NLMS being preferred because of its low computational complexity, stable convergence, and suitability for real-time implementation. Previous studies have also employed modulation-domain enhancement to reduce residual echo and background noise. However, most existing research evaluates adaptive filtering and speech enhancement separately or under limited acoustic conditions. This study contributes by combining these techniques into a single framework and evaluating it across multiple realistic environments.
The proposed methodology begins by generating room impulse responses to simulate acoustic echoes in different reverberant environments. The microphone signal is modeled as the combination of the desired near-end speech, acoustic echo, and background noise. The acoustic echo is produced by convolving the far-end speech with the room impulse response. The NLMS adaptive filter estimates the echo using the far-end speech as a reference and subtracts it from the microphone signal to produce a residual signal. The filter coefficients are continuously updated to ensure stable and efficient echo cancellation.
In the second stage, the residual signal is transformed into the time–frequency domain using the STFT. Noise characteristics are estimated, and a Wiener gain is computed from the estimated signal-to-noise ratio to suppress residual echo and background noise while preserving speech components. Finally, the enhanced speech signal is reconstructed using the ISTFT.
The system is implemented in MATLAB R2025b and evaluated under three simulated room environments with reverberation times (RT60) of 0.3 s, 0.6 s, and 0.9 s, six realistic background noise scenarios, and input signal-to-noise ratios (SNRs) ranging from −10 dB to 20 dB. Performance is assessed using objective metrics including Echo Return Loss Enhancement (ERLE), Output Signal-to-Noise Ratio (SNRout), Scale-Invariant Signal-to-Distortion Ratio (SI-SDR), and Short-Time Objective Intelligibility (STOI). These metrics measure echo suppression, speech enhancement quality, distortion reduction, and intelligibility, respectively.
Experimental analysis using waveforms and spectrograms demonstrates that the proposed framework effectively suppresses acoustic echoes across different room conditions. In smaller rooms with lower reverberation, the system produces clearer waveforms and better-preserved speech harmonics. Performance remains robust in more reverberant environments, indicating the framework’s ability to handle challenging acoustic conditions.
Conclusion
For hands-free communication systems, this work introduced a hybrid acoustic echo cancellation architecture that combines modulation-domain voice enhancement, NLMS adaptive filtering, and Room Impulse Response (RIR) simulation. Consistent increases in ERLE, output SNR, SI-SDR, and STOI were shown by experimental evaluation under three room environments, six representative background noise circumstances, and input SNR levels ranging from ?10 dB to 20 dB. Effective acoustic echo reduction and enhanced speech quality in various acoustic settings are confirmed by the results.
References
[1] E. P. Jayakumar, P. V. Muhammed Shifas, and P. S. Sathidevi, \"Integrated acoustic echo and noise suppression in modulation domain,\" International Journal of Speech Technology, vol. 19, no. 3, pp. 611–621, Sept. 2016.
[2] J. B. Allen and D. A. Berkley, \"Image method for efficiently simulating small-room acoustics,\" Journal of the Acoustical Society of America, vol. 65, no. 4, pp. 943–950, Apr. 1979.
[3] S. Haykin, Adaptive Filter Theory, 5th ed. Upper Saddle River, NJ, USA: Pearson Education, 2014.
[4] J. Benesty, T. Gänsler, D. R. Morgan, M. M. Sondhi, and S. L. Gay, Advances in Network and Acoustic Echo Cancellation. Berlin, Germany: Springer, 2001.
[5] P. C. Loizou, Speech Enhancement: Theory and Practice, 2nd ed. Boca Raton, FL, USA: CRC Press, 2013.
[6] Y. Ephraim and D. Malah, \"Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator,\" IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 32, no. 6, pp. 1109–1121, Dec. 1984.
[7] R. Martin, \"Noise power spectral density estimation based on optimal smoothing and minimum statistics,\" IEEE Transactions on Speech and Audio Processing, vol. 9, no. 5, pp. 504–512, Jul. 2001.
[8] K. K. Paliwal, K. Wójcicki, and B. Schwerin, \"Single-channel speech enhancement using spectral subtraction in the short-time modulation domain,\" Speech Communication, vol. 52, no. 5, pp. 450–475, 2010.
[9] K. K. Paliwal, B. Schwerin, and K. Wójcicki, \"Speech enhancement using a minimum mean-square error short-time spectral modulation magnitude estimator,\" Speech Communication, vol. 54, no. 2, pp. 282–305, Feb. 2012.
[10] MathWorks, \"Short-Time Fourier Transform (STFT),\" MathWorks Documentation.
[11] C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, \"A Short-Time Objective Intelligibility Measure for Time-Frequency Weighted Noisy Speech,\" IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2010.
[12] J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, \"SDR—Half-Baked or Well Done?\" IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019.